DRAFT — under teacher review.

Evaluation vs Testing — What's the Difference?

The Hamilton and Alexandra College · Year 12 · 2026

One of the most persistent mistakes in C4-2 is writing about testing when the rubric is asking about evaluation. These two processes are defined differently, use different evidence, and answer different questions. You can pass every test and still fail an evaluation criterion.


The one-sentence rule

Process One-sentence definition Question it answers
Testing Checking whether the software works correctly "Does it do what it is supposed to do?"
Evaluation Judging how well the software meets the criteria set at design time "Does it meet the standard we set?"

Testing is binary: the system either passes or fails a test case. Evaluation is a judgement against a criterion — and the criterion was derived from an SRS requirement.


A concrete example

Imagine a student builds an event-booking app. Here are three test results:

Test Result
User can search for events Pass
Booking confirmation email is sent Pass
All records saved correctly to file Pass

All tests pass. Now look at the evaluation matrix:

Criterion SRS Requirement Score (1–5) Result
A first-time user shall complete a booking in under 4 minutes without assistance NFR2 — Usability 2 Fails criterion
The event list shall load in under 2 seconds on school Wi-Fi NFR1 — Performance 4 Meets criterion

The app works (tests pass) but does not meet the usability standard (evaluation fails). Testing cannot catch this — only evaluation can.

This is the distinction the rubric language at 7–8 and 9–10 is checking:

"Uses the evaluation criteria to explain which elements of the design ideas should be further developed…" — 7–8 band

"Uses the evaluation criteria to justify which elements…" — 9–10 band

The rubric is asking you to reason from criterion scores, not from test results.


The most common mistakes

Mistake 1: Treating testing as evaluation

"I evaluated the app by running three test cases. All tests passed, so the design is good."

Test pass rates are not evaluation evidence. You need criterion scores linked to SRS requirements.

Mistake 2: Writing evaluation criteria that are just test cases

"Criterion: The search function returns results."

This is a test case disguised as a criterion — it is binary (pass/fail) and measures whether a feature exists, not how well it works. A genuine evaluation criterion must be measurable against a standard, for example: "The search function returns results in under 1.5 seconds on a standard device."

Mistake 3: Inventing criteria that float free of the SRS

"Criterion: The interface is attractive."

Attractiveness is a valid effectiveness factor, but only if your SRS included a requirement about it. If it is not in the SRS, you cannot evaluate against it — you would be judging the solution against a goal you never set.


Why VCAA cares

At 5–6, the rubric requires you to develop and apply evaluation criteria. This means you need actual criteria with scoring, not a description of test results.

At 7–8, you must use the evaluation criteria to explain decisions. This means referencing specific scores: "Design Idea A scored 4/5 for usability because..."

At 9–10, you must use the evaluation criteria to justify decisions. This means arguing that the scores demonstrate a design element is worth developing further, with explicit links back to SRS requirements.


See also


← Back to C04 Home · VCE Software Development Hub